Genome Biology
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match Genome Biology's content profile, based on 637 papers previously published here. The average preprint has a 0.47% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Daito, Y.; Uechi, M.; Kinoshita, T.; Tonosaki, K.
Show abstract
Background: Accurate identification of differentially methylated regions (DMRs) is fundamental to epigenomic research but remains challenging due to biological variability among replicates, heterogeneous effect sizes, and the tendency of adjacent cytosines to share similar methylation states. Many existing methods aggregate methylation measurements before statistical testing or do not explicitly account for replicate-level variability, contributing to elevated false-positive rates. Results: We developed glmmDMR, a DMR detection framework that combines generalized linear mixed models with a seed-based strategy for reconstructing DMRs from locally high-confidence signals while explicitly modeling replicate-level variability. Using simulated datasets with known ground-truth DMRs, we demonstrate that false-positive detections are more strongly associated with methylation variance among biological replicates than with the magnitude of methylation differences between groups. glmmDMR achieved higher precision than existing approaches while maintaining competitive recall, particularly for subtle methylation differences. Site-level modeling with beta regression provided the strongest overall performance, and seed-based region construction reduced artificial DMR fragmentation, improving recovery of true DMR boundaries and producing more contiguous, biologically interpretable DMRs. Applied to Arabidopsis thaliana ddm1 methylomes and a rice DEMETER-LIKE DNA demethylase mutant (Osdml3a-1), glmmDMR identified biologically meaningful DMRs, revealing widespread TE-associated hypomethylation and subtle TE-family-specific hypermethylation. Conclusions: Replicate-level methylation variance is an important determinant of DMR detection performance, and explicitly modeling this variance improves discrimination of biologically meaningful methylation changes from high-variance signals. By combining variance-aware statistical modeling with seed-based region construction, glmmDMR provides a robust framework for identifying contiguous, biologically interpretable DMRs across diverse methylome datasets.
Li, S.; Wang, Z.; Hu, Y.; Ni, Q.; Feng, C.; Hu, Y.; Zhang, S.; Chen, M.
Show abstract
Background3-tag-based sequencing methods have become the predominant approach for single-cell and spatial transcriptomics, with some protocols proven effective in detecting alternative polyadenylation (APA). While numerous computational tools have been developed for APA detection from these sequencing data, the absence of comprehensive benchmarks and the diversity of sequencing protocols and tools make it challenging to select appropriate methods for APA analysis in these contexts. ResultsWe systematically compared seven 3-tag-based sequencing protocols and identified key peak features affecting APA detection performance. We developed a simulation pipeline that generates realistic datasets preserving protocol-specific characteristics. Using simulated and real data, we comprehensively assessed six computational tools for their ability to identify polyA sites, quantify polyA site expression, detect differentially expressed (DE) APA genes, filter sequencing artifacts, and their computational efficiency. We also investigated factors influencing APA detection. Our evaluation revealed that SCAPE and scAPAtrap generally outperformed other tools across various performance metrics and protocols. ConclusionOur systematic evaluation provides guidance for tool selection, experiment design, and future tool development in APA analysis for singlecell and spatial transcriptomics, paving the way for investigating APA in these contexts.
Simmons, S. K.; Adiconis, X.; Haywood, N.; Parker, J.; Lin, Z.; Liao, Z.; Tuncali, I.; Al'Khafaji, A.; Shin, A.; Jagadeesh, K.; Gosik, K.; Gatzen, M.; Smith, J. T.; El Kodsi, D. N.; Kuras, Y.; Baecher-Allan, C.; Serrano, G. E.; Beach, T. G.; Garimella, K.; Rozenblatt-Rosen, O.; Regev, A.; Dong, X.; Scherzer, C.; Levin, J. Z.
Show abstract
Single-cell RNA-seq (scRNA-seq) is emerging as a powerful tool for understanding gene function across diverse cells. Recently, this has included the use of allele-specific expression (ASE) analysis to better understand how variation in the human genome affects RNA expression at the single-cell level. We reasoned that because intronic reads are more prevalent in single-nucleus RNA-Seq (snRNA-Seq), and introns are under lower purifying selection and thus enriched for genetic variants, that snRNA-seq should facilitate single-cell analysis of ASE. Here we demonstrate how experimental and computational choices can improve the results of allelic imbalance analysis. We explore how experimental choices, such as RNA source, read length, sequencing depth, genotyping, etc., impact the power of ASE-based methods. We developed a new suite of computational tools to process and analyze scRNA-seq and snRNA-seq for ASE. As hypothesized, we extracted more ASE information from reads in intronic regions than those in exonic regions and show how read length can be set to increase power. Additionally, hybrid selection improved our power to detect allelic imbalance in genes of interest. We also explored methods to recover allele-specific isoform expression levels from both long- and short-read snRNA-seq. To further investigate ASE in the context of human disease, we applied our methods to a Parkinsons disease cohort of 94 individuals and show that ASE analysis had more power than eQTL analysis to identify significant SNP/gene pairs in our direct comparison of the two methods. Overall, we provide an end-to-end experimental and computational approach for future studies.
Wissel, D.; Mehlferber, M. M.; Nguyen, K. M.; Pavelko, V.; Tseng, E.; Robinson, M. D.; Sheynkman, G. M.
Show abstract
PacBio long-read RNA sequencing resolves transcripts with greater clarity than short-read technologies, yet its quantitative performance remains under-evaluated at scale. Here, we benchmark the high-throughput PacBio Kinnex platform against Illumina short-read RNA-seq using matched, deeply sequenced datasets across a time course of endothelial cell differentiation. Compared to Illumina, Kin-nex achieved comparable gene-level quantification and more accurate transcript discovery and transcript quantification. While Illumina detected more transcripts overall, many reflected potentially unstable or ambiguous estimates in complex genes. Kinnex largely avoids these issues, producing more reliable differential transcript expression (DTE) calls, despite a mild bias against short transcripts (shorter than 1.25 kb). When correcting Illumina for inferential variability, Kinnex and Illumina quantifications were highly concordant, demonstrating equivalent performance. We also benchmarked long-read tools, nominating Oarfish as the most efficient for our Kinnex data. Together, our results establish Kinnex as a reliable platform for full-length transcript quantification.
Wynn, E. A.; Mould, K. J.; Vestal, B. E.; Moore, C. M.
Show abstract
Longitudinal scRNA-seq experiments offer a powerful approach for dissecting temporal gene expression dynamics in individual cell types. However, few methods have been developed specifically to address the unique statistical challenges of repeated measures in scRNA-seq data. Here, we introduce a novel method, REBEL (Repeated measures Empirical Bayes differential Expression analysis using Linear mixed models), for analyzing cell type-specific differential expression in repeated measures scRNA-seq experiments. Using simulation studies, we demonstrate that, relative to conventional repeated measures analysis methods and other scRNA-seq approaches, REBEL controls the false discovery rate and exhibits competitive power across a range of simulation scenarios. We further validate REBEL by analyzing a longitudinal scRNA-seq dataset from patients with B-cell lymphoma receiving chimeric antigen receptor (CAR)-T cell therapy. REBEL is implemented as an R package, available at https://github.com/ewynn610/REBEL.
Ostner, J.; Kirk, T.; Olayo-Alarcon, R.; Thöming, J. G.; Rosenthal, A. Z.; Häussler, S.; Müller, C. L.
Show abstract
Bacterial single-cell RNA sequencing has the potential to elucidate within-population heterogeneity of prokaryotes, as well as their interaction with host systems. Despite conceptual similarities, the statistical properties of bacterial single-cell datasets are highly dependent on the protocol, making proper processing essential to tap their full potential. We present BacSC, a fully data-driven computational pipeline that processes bacterial single-cell data without requiring manual intervention. BacSC performs data-adaptive quality control and variance stabilization, selects suitable parameters for dimension reduction, neighborhood embedding, and clustering, and provides false discovery rate control in differential gene expression testing. We validated BacSC on a broad selection of bacterial single-cell datasets spanning multiple protocols and species. Here, BacSC detected subpopulations in Klebsiella pneumoniae, found matching structures of Pseudomonas aeruginosa under regular and low-iron conditions, and better represented subpopulation dynamics of Bacillus subtilis. BacSC thus simplifies statistical processing of bacterial single-cell data and reduces the danger of incorrect processing.
Sigauke, R. F.; Sanford, L.; Maas, Z. L.; Jones, T.; Stanley, J. T.; Townsend, H. A.; Allen, M. A.; Dowell, R. D.
Show abstract
Gene transcription is controlled and modulated by regulatory regions, including enhancers and promoters. These regions are abundant in unstable, non-coding bidirectional transcription. Using nascent RNA transcription data across hundreds of human samples, we identified over 800,000 regions containing bidirectional transcription. We then identify highly correlated transcription between bidirectional and gene regions. The identified correlated pairs, a bidirectional region and a gene, are enriched for disease associated SNPs and often supported by independent 3D data. We present these resources as an SQL database which serves as a resource for future studies into gene regulation, enhancer associated RNAs, and transcription factors.
Stuart, T.; Srivastava, A.; Lareau, C.; Satija, R.
Show abstract
The recent development of experimental methods for measuring chromatin state at single-cell resolution has created a need for computational tools capable of analyzing these datasets. Here we developed Signac, a framework for the analysis of single-cell chromatin data, as an extension of the Seurat R toolkit for single-cell multimodal analysis. Signac enables an end-to-end analysis of single-cell chromatin data, including peak calling, quantification, quality control, dimension reduction, clustering, integration with single-cell gene expression datasets, DNA motif analysis, and interactive visualization. Furthermore, Signac facilitates the analysis of multimodal single-cell chromatin data, including datasets that co-assay DNA accessibility with gene expression, protein abundance, and mitochondrial genotype. We demonstrate scaling of the Signac framework to datasets containing over 700,000 cells. AvailabilityInstallation instructions, documentation, and tutorials are available at: https://satijalab.org/signac/
Hafemeister, C.; Halbritter, F.
Show abstract
Single-cell RNA sequencing (scRNA-seq) has become a standard approach to investigate molecular differences between cell states. Comparisons of bioinformatics methods for the count matrix transformation (normalization) and differential expression (DE) analysis of these data have already highlighted recommendations for effective between-sample comparisons and visualization. Here, we examine two remaining open questions: (i) What are the best combinations of data transformations and statistical test methods, and (ii) how do pseudo-bulk approaches perform in single-sample designs? We evaluated the performance of 343 DE pipelines (combinations of eight types of count matrix transformations and ten statistical tests) on simulated and real-world data, in terms of precision, sensitivity, and false discovery rate. We confirm superior performance of pseudo-bulk approaches without prior transformation. For within-sample comparisons, we advise the use of three pseudo-replicates, and provide a simple R package DElegate to facilitate application of this approach.
Wang, C.; Prawer, Y. D. J.; Voogd, O.; Schuster, J.; Pasquali, C.; De Paoli-Iseppi, R.; Li, A.; Hallab, J.; Tian, L.; Peng, H.; David, M.; Du, M. R. M.; Velasco, S.; Garone, M. G.; Dong, X.; Zeglinski, K.; Pavan, C.; Law, K. C. L.; Abu-Bonsrah, K. D.; Hunt, C. P. J.; Parish, C.; Gouil, Q.; Thijssen, R.; Davidson, N. M.; Ritchie, M. E.; Clark, M. B.; You, Y.
Show abstract
Long-read single-cell RNA-sequencing enables the profiling of RNA isoform expression and alternative splicing at single cell resolution. However, diverse single-cell technologies and sparse isoform data demand flexible and accurate analysis tools. We introduce FLAMESv2, a highly modular and protocol-agnostic R/Bioconductor package for long-read single-cell RNA-seq data analysis. FLAMESv2 supports a wide range of single-cell and spatial protocols, is highly configurable, scales to allow multi-sample analysis and provides versatile visualisation and analysis outputs. We demonstrate its compatibility with both droplet-based and combinatorial barcoding single-cell methods, as well as spatial transcriptomics workflows. Benchmarking confirms FLAMESv2 achieves field-leading performance across key analysis tasks. Applying FLAMESv2 to in vitro differentiation of stem cells into neurons, we identify cell-types, differentiation trajectories, expression of annotated and novel isoforms and isoform expression diversity and heterogeneity within individual cells. FLAMESv2 provides a comprehensive, flexible approach to analysing long-read single-cell RNA-sequencing, unlocking this powerful methodology for RNA isoform characterisation.
Ardaman, A.; Forgiarini, C.; Arunkumar, R.
Show abstract
Intraspecific hybridization in allopolyploid plant genomes has the potential to induce non-additive changes in gene expression and DNA cytosine methylation, partly through interactions among divergent parental subgenomes. However, the extent to which intraspecific hybridization reshapes gene expression, coordinates homoeolog regulation, and remodels methylation in higher-order polyploids remains poorly quantified. To address this, we sequenced seedling leaf transcriptomes and methylomes from two parental cultivars of hexaploid bread wheat (Triticum aestivum L.) and their hybrids. More than 40% of genes were differentially expressed between hybrids and parents, although many were not differentially expressed between the parents themselves, consistent with complex trans-regulatory effects in the hybrid genome. This effect was more pronounced for homoeologs whose relative expression differed between the parents. These expression shifts often occurred simultaneously across all three homoeologs within triads, reducing homoeolog expression bias (HEB) in the hybrids. CG methylation levels were similar between the parents and hybrids in regions of low genetic divergence and in transposable element (TE)-rich regions, whereas CG sites in gene-rich regions showed more additive inheritance (hybrids intermediate between parents), particularly when parental haplotypes were themselves divergent. TE and gene body methylation (gbM) was strongly conserved in parents and hybrids. gbM was associated with more balanced homoeolog expression and fewer non-additive expression changes. CHH methylation showed overdominance, whereas non-conserved CHG methylation was enriched in TE-rich regions, suggesting that non-CG remodeling may reflect parental differences in TE and small-RNA content. Our results show that intraspecific hybridization within a hexaploid species can generate non-additive changes in gene expression and DNA methylation in seedling leaf tissue, while the presence of homoeologous genes, parental HEB, parental genetic and methylation divergence, and genomic location have varying levels of influence on expression or methylation remodeling.
ZOU, B.; Wang, J.; Ding, Y.; Zhang, Z.; Yufen, H.; Fang, X.; Cheung, K. C.; See, S.; Zhang, L.
Show abstract
Metagenome-assembled genomes (MAGs) offer valuable insights into the exploration of microbial dark matter using metagenomic sequencing data. However, there is a growing concern that contamination in MAGs may significantly impact the downstream analysis results. Existing MAG decontamination methods heavily rely on marker genes but do not fully leverage genomic sequences. To address the limitations, we have introduced a novel decontamination approach named Deepurify, which utilizes a multi-modal deep language model employing contrastive learning to learn taxonomic similarities of genomic sequences. Deepurify utilizes inferred taxonomic lineages to guide the allocation of contigs into a MAG-separated tree and employs a tree traversal strategy for maximizing the total number of medium- and high-quality MAGs. Extensive experiments were conducted on two simulated datasets, CAMI I, and human gut metagenomic sequencing data. These results demonstrate that Deepurify significantly outperforms other decontamination methods.
Crowell, H. L.; Leonardo, S. X. M.; Soneson, C.; Robinson, M. D.
Show abstract
With the emergence of hundreds of single-cell RNA-sequencing (scRNA-seq) datasets, the number of computational tools to analyse aspects of the generated data has grown rapidly. As a result, there is a recurring need to demonstrate whether newly developed methods are truly performant - on their own as well as in comparison to existing tools. Benchmark studies aim to consolidate the space of available methods for a given task, and often use simulated data that provide a ground truth for evaluations. Thus, demanding a high quality standard for synthetically generated data is critical to make simulation study results credible and transferable to real data. Here, we evaluated methods for synthetic scRNA-seq data generation in their ability to mimic experimental data. Besides comparing gene- and cell-level quality control summaries in both one- and two-dimensional settings, we further quantified these at the batch- and cluster-level. Secondly, we investigate the effect of simulators on clustering and batch correction method comparisons, and, thirdly, which and to what extent quality control summaries can capture reference-simulation similarity. Our results suggest that most simulators are unable to accommodate complex designs without introducing artificial effects; they yield over-optimistic performance of integration, and potentially unreliable ranking of clustering methods; and, it is generally unknown which summaries are important to ensure effective simulation-based method comparisons.
Groth, T. E.; Mishin, A. A.; Rao, V.; Tibet, R.; Troll, C. J.
Show abstract
Cell-free DNA methylation sequencing provides insight into tissue of origin and chromatin structure. In some workflows, generating libraries includes end-repair. Using matched single-stranded and double-stranded libraries prepared from the same cfDNA extracts, we show that end-repair in double-stranded DNA libraries reduces globally inferred CpG methylation leading to decreased tissue of origin accuracy. Trimming read termini partially mitigates this bias but decreases coverage and removes fragmentomic information compared to single-stranded DNA libraries, which forego end-repair.
Li, Y.; Wang, T.-Y.; Guo, Q.; Ren, Y.; Lu, X.; Cao, Q.; Yang, R.
Show abstract
Chimera artifacts in nanopore direct RNA sequencing (dRNA-seq) can significantly distort transcriptome analyses, yet their detection and removal remain challenging due to limitations in existing basecalling models. We present Deep-Chopper, a genomic language model that precisely identifies and removes adapter sequences from base-called dRNA-seq long reads at single-base resolution, operating independently of raw signal or alignment information to effectively eliminate chimeric read artifacts. By removing these artifacts, DeepChopper substantially improves the accuracy of critical downstream analyses, such as transcript annotation and gene fusion detection, thereby enhancing the reliability and utility of nanopore dRNA-seq for transcriptomics research.
Ratnasiri, K.; Mach, S. N.; Blish, C. A.; Khatri, P.
Show abstract
Traditional differential gene expression methods are limited for analysis of single cell RNA-sequencing (scRNA-seq) studies that use paired repeated measures and matched cohort designs. Many existing approaches consider cells as independent samples, leading to high false positive rates while ignoring inherent sampling structures. Although pseudobulk methods address this, they ignore intra-sample expression variability and have higher false negatives rates. We propose a novel meta-analysis approach that accounts for biological replicates and cell variability in paired scRNA-seq data. Using both real and synthetic datasets, we show that our method, single-cell MetaIntegrator (https://github.com/Khatri-Lab/scMetaIntegrator), provides robust effect size estimates and reproducible p-values.
Reiter, T. E.; Pierce-Ward, N. T.; Irber, L. C.; Botvinnik, O.; Brown, C. T.
Show abstract
An estimated 2 billion species of microbes exist on Earth with orders of magnitude more strains. Microbial pangenomes are created by aggregating all genomes of a single clade and reflect the metabolic diversity of groups of organisms. As de novo metagenome analysis techniques have matured and reference genome databases have expanded, metapangenome analysis has risen in popularity as a tool to organize the functional potential of organisms in relation to the environment from which those organisms were sampled. However, the reliance on assembly and binning or on reference databases often leaves substantial portions of metagenomes unanalyzed, thereby underestimating the functional potential of a community. To address this challenge, we present a method for metapangenomics that relies on amino acid k-mers (kaa-mers) and metagenome assembly graph queries. To enable this method, we first show that kaa-mers estimate pangenome characteristics and that open reading frames can be accurately predicted from short shotgun sequencing reads using the previously developed tool orpheum. These techniques enable pangenomics to be performed directly on short sequencing reads. To enable metapangenome analysis, we combine these approaches with compact de Bruijn assembly graph queries to directly generate sets of sequencing reads for a specific species from a metagenome. When applied to stool metagenomes from an individual receiving antibiotics over time, we show that these approaches identify strain fluctuations that coincide with antibiotic exposure.
Hu, Y.; Liu, Z.; Tsao, D.; Leung, J. M.; V. Gerayeli, F.; Li, X.; Shao, X.; Sin, D.; Zhang, X.
Show abstract
BackgroundPseudo-bulk RNA-seq, generated by aggregating single-cell profiles, is widely used for benchmarking deconvolution methods because it globally approximates bulk transcriptomes and provides known cell-type proportions as ground truth. However, pseudo-bulk inherits single-cell specific measurement properties, and the extent to which these differ from real bulk RNA-seq remains difficult to quantify in the absence of paired data. In practice, deconvolution studies also commonly rely on large external single-cell references, yet the value of small, protocol-matched in-study references has not been systematically evaluated under real bulk conditions. These gaps motivate a paired benchmark that jointly examines pseudo-bulk fidelity, reference design, and their consequences for deconvolution accuracy. ResultsWe establish a fully paired benchmarking framework using split-sample, donor-matched bulk and single-cell RNA-seq (scRNA-seq) from human bronchoalveolar lavage (BAL). Embedded within a reverse five-fold cross-validation design and evaluated on real bulk RNA-seq data, this framework benchmarks 15 deconvolution algorithms published in 2013-2025 across 3 cell-type resolutions and 5 single-cell references, including three published BAL datasets, a harmonized lung BAL atlas, and an in-study reference derived from paired aliquots. We show that real bulk and matched pseudo-bulk profiles exhibit systematic gene-level differences, identifying 557 reproducibly discordant genes (|log2 FC| > 1, FDR< 0.05 by LIMMA), including cell-type informative features. These discrepancies reflect technology-specific effects and violate the linear mixing assumption underlying most deconvolution methods. We demonstrate that a protocol-matched in-study reference constructed from only six donors consistently outperforms substantially larger external references, with the advantage that increases at finer cell-type resolution. Moreover, paired samples enable the identification and selective removal of discordant genes, which further improves deconvolution accuracy for many algorithms, particularly in high-resolution settings. These findings are robust across methods, references, and evaluation criteria and extend beyond compositional accuracy to improving recovery of disease-associated cell-type differences in clinical applications. ConclusionsOur study provides the first fully paired benchmark of transcriptomic deconvolution on real bulk RNA-seq data of human BAL samples and demonstrates that reference design and data compatibility are as influential as algorithm choice. Beyond benchmarking, we introduce a practical and cost-effective protocol for deconvolution studies: generate single-cell data for a minimal subset of bulk samples (pilot pairing), use these data to construct an in-study reference and identify discordant genes, and apply the resulting insights to the full cohort. This strategy requires limited additional experimental effort yet yields substantial gains in accuracy and stability, offering actionable guidance for future bulk RNA-seq deconvolution studies across tissues and platforms. One-sentence summarySystematic benchmarking reveals pseudo-bulk biases and provides practical fixes for accurate RNA deconvolution.
Duitama Gonzalez, C.; Lopopolo, M.; Nishimura, L.; Faure, R.; Duchene, S.
Show abstract
The field of ancient metagenomics provides insights into past microbiomes, but with a growing dataset size, methods that rely on reference databases have limited scope. Here, we introduce DIANA, a multi-task neural network that predicts key metadata categories from unitig abundances. Trained on 2,597 run accessions (1.72 Tbp of assembled unitig sequences), DIANA accurately identifies sample host (94.6%), community type (90.0%), and material (88.9%) on held-out test data and demonstrates robust generalisation on an independent validation set. A key innovation is DIANAs ability to perform semantic generalisation, correctly classifying samples with labels unseen during training -- such as novel subspecies -- to their appropriate parent categories. By leveraging both known and uncharacterized genomic sequences, DIANA provides a rapid, data-driven system for metadata validation and quality control, accelerating discovery in ancient metagenomics research.
Daneshpajouh, H.; Moghul, I.; Wiese, K. C.; Libbrecht, M. W.
Show abstract
The International Human Epigenome Consortium has generated thousands of epigenomic datasets that mea-sure various biochemical activities in the genome, including transcription factor binding, histone modification, and DNA accessibility. Currently, the predominant methods for integrating these datasets to annotate regu-latory elements are segmentation and genome annotation (SAGA) algorithms. The majority of annotations by these methods are cell type-specific. However, as the number of profiled cell types has grown into the thousands, using thousands of cell type-specific chromatin state annotations proves undesirable for many applications. Here, we present a pan-cell type annotation that summarizes all IHEC epigenomes using the recently-developed method, epigenome-ssm.